Papers with coarse-to-fine attention process
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Multimodal large language models may deviate from this pattern due to attention drift and underutilization of visual evidence. |
| Approach: | They propose a Dual-Indicator Guided Contrastive Alignment (DICA) that tracks visual attention and output image correlations to improve visual grounding. |
| Outcome: | The proposed model outperforms existing approaches and significantly reduces hallucinations. |